Search CORE

9 research outputs found

SIGMORPHON 2021 Shared Task on Morphological Reinflection: Generalization Across Languages

Author: Aiton Grant
Ambridge Ben
Ataman Duygu
Ate Yustinus Ghanggo
Barta Botond
Bayyr-ool Aziyana
Bernardy Jean-Philippe
Chodroff Eleanor
Coler Matt
Cotterell Ryan
Ek Adam
El-Khaissi Charbel
Ganieva Sofya
Gasser Michael
Goldman Omer
Habash Nizar
Hatcher Richard J.
Hulden Mans
Ivanova Sardana
Khalifa Salam
Kieraś Witold
Klyachko Elena
Krizhanovsky Andrew
Krizhanovsky Natalia
Kumar Ritesh
Lakatos Dorina
Lane William
Leonard Brian
Liu Zoey
Mielke Sabrina J.
Montoya Samame Jaime Rafael
Nicolai Garett
Nuriah Zahroh
Oncevay Arturo
Pimentel Tiago
Plugaryov Matvey
Ponti Edoardo M.
Prud'hommeaux Emily
Raj Mohit
Ratan Shyam
Ryskina Maria
Salchak Aelita
Salehi Ali
Shcherbakov Andrey
Sheifer Karina
Silva Villegas Gema Celeste
Stoehr Niklas
Straughn Christopher
Suhardijanto Totok
Szolnok Gábor
Tyers Francis M.
Vania Clara
Vylomova Ekaterina
Washington Jonathan
Woliński Marcin
Wu Shijie
Yarowsky David
Ács Judit
Publication venue: The Association for Computational Linguistics
Publication date: 01/08/2021
Field of study

This year's iteration of the SIGMORPHON Shared Task on morphological reinflection focuses on typological diversity and cross-lingual variation of morphosyntactic features. In terms of the task, we enrich UniMorph with new data for 32 languages from 13 language families, with most of them being under-resourced: Kunwinjku, Classical Syriac, Arabic (Modern Standard, Egyptian, Gulf), Hebrew, Amharic, Aymara, Magahi, Braj, Kurdish (Central, Northern, Southern), Polish, Karelian, Livvi, Ludic, Veps, Võro, Evenki, Xibe, Tuvan, Sakha, Turkish, Indonesian, Kodi, Seneca, Asháninka, Yanesha, Chukchi, Itelmen, Eibela. We evaluate six systems on the new data and conduct an extensive error analysis of the systems' predictions. Transformer-based models generally demonstrate superior performance on the majority of languages, achieving >90% accuracy on 65% of them. The languages on which systems yielded low accuracy are mainly under-resourced, with a limited amount of data. Most errors made by the systems are due to allomorphy, honorificity, and form variation. In addition, we observe that systems especially struggle to inflect multiword lemmas. The systems also produce misspelled forms or end up in repetitive loops (e.g., RNN-based models). Finally, we report a large drop in systems' performance on previously unseen lemmas.Peer reviewe

Edinburgh Research Explorer

Helsingin yliopiston digitaalinen arkisto

The Universal Morphology (UniMorph) project is a collaborative effort providing broad-coverage instantiated normalized morphological inflection tables for hundreds of diverse world languages. The project comprises two major thrusts: a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. This paper presents the expansions and improvements made on several fronts over the last couple of years (since McCarthy et al. (2020)). Collaborative efforts by numerous linguists have added 67 new languages, including 30 endangered languages. We have implemented several improvements to the extraction pipeline to tackle some issues, e.g. missing gender and macron information. We have also amended the schema to use a hierarchical structure that is needed for morphological phenomena like multiple-argument agreement and case stacking, while adding some missing morphological features to make the schema more inclusive. In light of the last UniMorph release, we also augmented the database with morpheme segmentation for 16 languages. Lastly, this new release makes a push towards inclusion of derivational morphology in UniMorph by enriching the data and annotation schema with instances representing derivational processes from MorphyNet

Proceedings - University of Groningen

University of Groningen

ARTS repository - University of Groningen

Dissertations of the University of Groningen

UniMorph 4.0: Universal Morphology

Author: Batsuren Khuyagbaatar
Bella Gábor
Budianskaya Elena
Coler Matt
Cotterell Ryan
El-Khaissi Charbel
et al.
Gasser Michael
Ghanggo Ate Yustinus
Goldman Omer
Gorman Kyle
Habash Nizar
Khalifa Salam
Kieraś Witold
Lane William Abbott
Leonard Brian
Mielke Sabrina
Montoya Samame Jaime Rafael
Nicolai Garrett
Pimentel Tiago
Raj Mohit
Ryskina Maria
Stöhr Niklas Werner
Publication venue: European Language Resources Association
Publication date: 01/01/2022
Field of study

Repository for Publications and Research Data

SIGMORPHON–UniMorph 2022 Shared Task 0: Generalization and Typologically Diverse Morphological Inflection

Author: Akkuş Faruk
Anastasopoulos Antonios
Andrushko Taras
Arora Aryaman
Atanalov Nona
Batsuren Khuyagbaatar
Bella Gábor
Budianskaya Elena
Cotterell Ryan
Dolatian Hossep
Ghanggo Ate Yustinus
Goldman Omer
Guriel David
Guriel Simon
Guriel-Agiashvili Silvia
Khalifa Salam
Kieraś Witold
Kodner Jordan
Krizhanovsky Andrew
Krizhanovsky Natalia
Marchenko Igor
Markowska Magdalena
Mashkovtseva Polina
Nepommiashchaya Maria
Rodionova Daria
Serova Alexandra
Sheifer Karina
Vylomova Ekaterina
Yemelina Anastasia
Young Jeremiah
Publication venue: Association for Computational Linguistics
Publication date: 01/07/2022
Field of study

The 2022 SIGMORPHON–UniMorph shared task on large scale morphological inflection generation included a wide range of typologically diverse languages: 33 languages from 11 top-level language families: Arabic (Modern Standard), Assamese, Braj, Chukchi, Eastern Armenian, Evenki, Georgian, Gothic, Gujarati, Hebrew, Hungarian, Itelmen, Karelian, Kazakh, Ket, Khalkha Mongolian, Kholosi, Korean, Lamahalot, Low German, Ludic, Magahi, Middle Low German, Old English, Old High German, Old Norse, Polish, Pomak, Slovak, Turkish, Upper Sorbian, Veps, and Xibe. We emphasize generalization along different dimensions this year by evaluating test items with unseen lemmas and unseen features separately under small and large training conditions. Across the five submitted systems and two baselines, the prediction of inflections with unseen features proved challenging, with average performance decreased substantially from last year. This was true even for languages for which the forms were in principle predictable, which suggests that further work is needed in designing systems that capture the various types of generalization required for the world’s languages

Repository for Publications and Research Data

Recommended from our members

Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss.

Author: Abbas Noor Karolin
Atkinson Quentin D
Auer Daniel
Bakker Nancy A
Barbos Giulia
Blasi Damián E
Borges Robert D
Bowern Claire
Chira Angela
Collins Jeremy
Danielsen Swintha
de Sousa Hilário
Dinnage Russell
Dorenbusch Luise
Dorn Ella
Dunn Michael
Elliott John
Epps Patience
Evans Nicholas
Falcone Giada
Fischer Jana
Forkel Robert
Ghanggo Ate Yustinus
Gibson Hannah
Goodall Jemima A
Gray Russell D
Greenhill Simon J
Gruner Victoria
Göbel Hans-Philipp
Hammarström Harald
Harvey Andrew
Haspelmath Martin
Hayes Rebekah
Haynie Hannah J
Heer Leonard
Herrera Miranda Roberto E
Hill Jane
Huntington-Rainey Biu
Hübler Nataliia
Ivani Jessica K
Johns Marilen
Just Erika
Kashima Eri
Kipf Carolina
Klingenberg Janina V
Koti Aikaterina
Kowalik Richard GA
Krasnoukhova Olga
König Nikita
Latarche Jay J
Lesage Jakob
Levinson Stephen C
Lindvall Nora LM
Lorenzen Mandy
Lutzenberger Hannah
Martins Tânia RA
Mata German Celia
Maurits Luke
Montoya Samamé Jaime
Muradoglu Saliha
Müller Michael
Neely Kelsey
Nickel Johanna
Norvik Miina
Oluoch Cheryl Akinyi
Passmore Sam
Peacock Jesse
Pearey India OC
Peck Naomi
Petit Stephanie
Pieper Sören
Poblete Mariana
Prestipino Daniel
Raabe Linda
Raja Amna
Reesink Ger
Reimringer Janis
Rey Sydney C
Rizaew Julia
Robbeets Martine
Ruppert Eloisa
Salmon Kim K
Sammet Jill
Schembri Rhiannon
Schlabbach Lars
Schmidt Frederick WP
Singer Ruth
Skilton Amalia
Skirgård Hedvig
Smith Wikaliler Daniel
Sverredal Kristin
Valle Daniel
van der Meer Suzanne
Vera Javier
Vesakoski Outi
Voß Judith
Weber Tobias
Witte Tim
Witzlack-Makarevich Alena
Wu Henry
Yam Stephanie
Ye Jingting
Yong Maisie
Yuditha Tessa
Zariquiey Roberto
Publication venue: Sci Adv
Publication date: 02/05/2023
Field of study

While global patterns of human genetic diversity are increasingly well characterized, the diversity of human languages remains less systematically described. Here, we outline the Grambank database. With over 400,000 data points and 2400 languages, Grambank is the largest comparative grammatical database available. The comprehensiveness of Grambank allows us to quantify the relative effects of genealogical inheritance and geographic proximity on the structural diversity of the world's languages, evaluate constraints on linguistic diversity, and identify the world's most unusual languages. An analysis of the consequences of language loss reveals that the reduction in diversity will be strikingly uneven across the major linguistic regions of the world. Without sustained efforts to document and revitalize endangered languages, our linguistic window into human history, cognition, and culture will be seriously fragmented

Apollo (Cambridge)

Grambank reveals the importance of genealogical constraints on linguistic diversity and highlights the impact of language loss

Author: Abbas Noor Karolin
Atkinson Quentin D.
Auer Daniel
Bakker Nancy A.
Barbos Giulia
Blasi Dami\ue1n E.
Borges Robert D.
Bowern Claire
Chira Angela
Collins Jeremy
Danielsen Swintha
de Sousa Hil\ue1rio
Dinnage Russell
Dorenbusch Luise
Dorn Ella
Dunn Michael
Elliott John
Epps Patience
Evans Nicholas
Falcone Giada
Fischer Jana
Forkel Robert
G\uf6bel Hans-Philipp
Ghanggo Ate Yustinus
Gibson Hannah
Goodall Jemima A.
Gray Russell D.
Greenhill Simon J.
Gruner Victoria
H\ufcbler Nataliia
Hammarstr\uf6m Harald
Harvey Andrew
Haspelmath Martin
Hayes Rebekah
Haynie Hannah J.
Heer Leonard
Herrera Miranda Roberto E.
Hill Jane
Huntington-Rainey Biu
Ivani Jessica K.
Johns Marilen
Just Erika
K\uf6nig Nikita
Kashima Eri
Kipf Carolina
Klingenberg Janina V.
Koti Aikaterina
Kowalik Richard G.\ua0A.
Krasnoukhova Olga
Latarche Jay J.
Lesage Jakob
Levinson Stephen C.
Lindvall Nora L.\ua0M.
Lorenzen Mandy
Lutzenberger Hannah
M\ufcller Michael
Martins T\ue2nia R.\ua0A.
Mata German Celia
Maurits Luke
Montoya Samam\ue9 Jaime
Muradoglu Saliha
Neely Kelsey
Nickel Johanna
Norvik Miina
Oluoch Cheryl Akinyi
Passmore Sam
Peacock Jesse
Pearey India O.\ua0C.
Peck Naomi
Petit Stephanie
Pieper S\uf6ren
Poblete Mariana
Prestipino Daniel
Raabe Linda
Raja Amna
Reesink Ger
Reimringer Janis
Rey Sydney C.
Rizaew Julia
Robbeets Martine
Ruppert Eloisa
Salmon Kim K.
Sammet Jill
Schembri Rhiannon
Schlabbach Lars
Schmidt Frederick W.\ua0P.
Singer Ruth
Skilton Amalia
Skirg\ue5rd Hedvig
Smith Wikaliler Daniel
Sverredal Kristin
Valle Daniel
van der Meer Suzanne
Vera Javier
Vesakoski Outi
Vo f Judith
Weber Tobias
Witte Tim
Witzlack-Makarevich Alena
Wu Henry
Yam Stephanie
Ye Jingting
Yong Maisie
Yuditha Tessa
Zariquiey Roberto
Publication venue
Publication date: 01/01/2023
Field of study

Institutional Repository Universiteit Antwerpen